Papers with Concadia dataset
Updating CLIP to Prefer Descriptions Over Captions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current metrics for imagetext similarity tend to be insensitive to the text's purpose. |
| Approach: | They propose to use a model that assigns higher scores to descriptions than captions . they use parameter efficient fine-tuning and a loss objective to shed light on the distinction . |
| Outcome: | The proposed model correlates with the judgements of blind and low-vision people while preserving transfer capabilities and sheds light on the caption–description distinction. |